Papers with audio quality

20 papers
Towards Codec-LM Co-design for Neural Codec Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Neural codec language models (or codec LMs) are emerging as a powerful framework for text-to-speech (TTS) despite the close interdependence of codecs and LM, research on codec and lms has largely remained siloed.
Approach: They propose a frame-wise codec encoder that improves both LM log-likelihood and TTS metrics . they also propose LM codebook level dropout to efficiently navigate a portion of codec-LM design space .
Outcome: The proposed codec-LM co-design improves intelligibility, audio quality and speaker control compared to a siloed baseline.
On the Semantic Latent Space of Diffusion-Based Text-To-Speech Models (2024.acl-short)

Copied to clipboard

Challenge: Denoising Diffusion Models (DDMs) are a powerful generative tool for text-to-speech (TTS) but their semantic capabilities are unknown and control of synthesized speech’s vocal properties remains a challenge.
Approach: They explore the latent space of frozen TTS models composed of latent bottleneck activations of the DDM’s denoiser and propose methods for finding semantic directions within it.
Outcome: The proposed methods enable off-the-shelf audio editing without any training, architectural changes or data requirements.
Multimodal Generation with Consistency Transferring (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for multimodal content generation are limited to unimodal content production due to high training complexity, significant costs, and inadequate emphasis on model constraints.
Approach: They propose a method to generate multimodal content with constraints on adjacent steps and a layer-based layer-constrained transfer between adjacent steps to improve denoising capabilities.
Outcome: The proposed method improves the model’s ability to capture actions and depict backgrounds more effectively and improves video generation speed by approximately 40% and quality by about 39.3%.
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching (2025.naacl-long)

Copied to clipboard

Challenge: Flow matching is a robust and stable approach to training diffusion models, but it can result in subpar audio quality.
Approach: They propose a reparameterized flow matching model for mel-spectrogram conditioned speech synthesis that uses a mel prior instead of a standard Gaussian prior to minimize unnecessary transportation costs.
Outcome: The proposed model improves sample quality and generation speed for speech vocoders while reducing transportation costs.
The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)

Copied to clipboard

Challenge: The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody .
Approach: They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language .
Outcome: The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies.
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt (2024.naacl-long)

Copied to clipboard

Challenge: Recent singing-voice-synthesis methods lack ability to control style attributes of synthesized singing.
Approach: They propose a singing-voice-synthesis method that enables attribute controlling on singer gender, vocal range and volume with natural language.
Outcome: The proposed method achieves favorable control ability and audio quality.
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control (2025.acl-long)

Copied to clipboard

Challenge: Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS systems cannot perform speaker-specific voice generation.
Approach: They propose a style control module that captures codec representations corresponding to timbre, content, and style in a discrete decoupling codec space.
Outcome: The proposed system can fully clone the speaker's voice and perform speech-specific adjustment and control functions.
Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)

Copied to clipboard

Challenge: Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation .
Approach: They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space.
Outcome: The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis.
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers (2024.emnlp-main)

Copied to clipboard

Challenge: Existing neural speech codecs trade model complexity for reconstruction performance . ESC is a lightweight, parameter-efficient speech coder .
Approach: They propose an efficient speech codec based on a cross-scale residual vector quantization scheme and transformers that can achieve high-fidelity speech reconstruction with significantly lower model complexity.
Outcome: The proposed codec achieves high-fidelity speech reconstruction with significantly lower model complexity.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been used for general-purpose interfaces across multiple tasks and languages.
Approach: They propose to use large language models as a general-purpose interface across multiple tasks and languages.
Outcome: The proposed model performs better on 200K hours of 6-language data for voice generation applications.
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment (2025.acl-long)

Copied to clipboard

Challenge: Existing zero-shot text-to-speech systems struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis.
Approach: They propose a dataset that leverages preference alignment techniques to improve performance . they also extend the Direct Preference Optimization framework to accommodate diverse TTS architectures .
Outcome: The proposed dataset improves intelligibility, similarity, and audio quality for multiple models across domains.
Non-parallel Accent Transfer based on Fine-grained Controllable Accent Modelling (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing accent transfer methods rely on parallel data or speech recognition models.
Approach: They propose to use mutual information learning to disentangle accent features and control the accent of the generated speech during the inference time.
Outcome: The proposed framework achieves superior performance to baseline models in accentedness and audio quality.
FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment.
Approach: They propose to learn straight flow for fast simulation by using flashAudio with rectified flows and immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment.
Outcome: The proposed method can learn straight flow for fast simulations and reduce noise distribution.
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously.
Approach: They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has.
Outcome: The proposed method beats text-based systems in terms of perceived emotion and audio quality.
Call My Net 2: A New Resource for Speaker Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Call My Net 2 (CMN2) corpus features Tunisian Arabic conversations between friends and family . call recordings include speech in various realistic and natural acoustic settings, both noisy and non-noisy.
Approach: They introduce the Call My Net 2 (CMN2) corpus, a new resource for speaker recognition featuring Tunisian Arabic conversations between friends and family.
Outcome: The Call My Net 2 (CMN2) corpus contains data from over 400 Tunisian Arabic speakers . each speaker made 10 or more calls each lasting up to 10 minutes .
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios is a major challenge for end-to end spoken dialogue models.
Approach: They propose to provide an extensive evaluation framework for end-to-end spoken dialogue models (SDMs) that includes both cognitive dimensions and paralinguistic cues .
Outcome: The proposed benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model’s abilities in U**nderstanding, **R**easoning, and **O**ral conversation.
Towards a Corpus of Spoken Maltese: Korpus tal-Malti Mitkellem, KMM (2024.lrec-main)

Copied to clipboard

Challenge: 'Corpus of Spoken Maltese' is a spoken corpus of spoken Malteser based on a gold-standard Core collection . initial results show that the ASR is robust enough to generate first-pass texts for annotators to work on, thus reducing the human effort and consequently, the cost involved.
Approach: They propose to create a “dedicated” spoken corpus of Maltese based on a gold-standard Core collection and a qualitative analysis of the output of a Malteser ASR system.
Outcome: The proposed corpus is based on the concept of a gold-standard Core collection and compares to human annotations.
Hierarchical Representation Alignment Learning of Diffusion Transformers for Neural Audio Codec (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in diffusion and conditional flow matching models for low-resolution domains are underexplored.
Approach: They propose a CFM-based model that iteratively generates raw waveform in low-bitrate conditions . they propose DVQ, a factorized quantization method that uses a single quantizer .
Outcome: The proposed model outperforms state-of-the-art neural audio codecs in audio quality and semantic intelligibility under low-bitrate conditions.
FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in speech synthesis have improved audio quality and pronunciation . fillers are an integral part of natural human conversation, but achieving human-like conversational speech remains a challenge.
Approach: They propose a speech synthesis framework that enables natural filler insertion and style control . they propose 'filler-inclusive' speech data that includes fillers with pitch and duration information .
Outcome: The proposed framework enables natural filler insertion and style control . the proposed framework is validated and can be used to predict filler style .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations